Papers with answer-free policy

1 papers
AG-GRPO: Answer-Guided GRPO for Masked Diffusion Language Models (2026.acl-long)

Copied to clipboard

Challenge: Recent work on large language models (LLMs) has emphasized not only final-answer accuracy but also reliability of reasoning on challenging tasks.
Approach: They propose an answer-guided group-relative policy optimization for masked diffusion language models which generates text through iterative mangled token restoration.
Outcome: The proposed approach improves over pretrained dLLMs and prior RL methods across mathematics, puzzle-solving, and code-generation benchmarks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations